Papers by Aditya K Surikuchi

    1 papers
    LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

    Copied to clipboard

    Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
    Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
    Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations